Same Data, Different Answers: The Hidden Sources of Irreproducibility in AI Training
When two research teams train identical models on identical datasets and arrive at meaningfully different results, the scientific integrity of the entire enterprise comes into question. This investigation examines the technical and procedural fault lines—from floating-point arithmetic to GPU-level variance—that make reproducibility in modern AI research far more elusive than the field typically acknowledges. Understanding these sources of non-determinism is not merely an academic exercise; it is